Skip to main content

61. LLM Optimization

Important optimization techniques include:

  • quantization
  • batching
  • KV caching
  • efficient attention
  • optimized kernels
  • speculative decoding

The goal can be:

lower latency
higher throughput
lower memory
lower cost

Optimization requires measuring the actual bottleneck rather than changing components blindly.